Papers with social reasoning

9 papers
Social Intelligence in the Age of LLMs (2025.naacl-tutorial)

Copied to clipboard

Challenge: Large Language Models (LLMs) are a powerful tool for integrating human-like communication and context-aware interactions into artificial systems.
Approach: They propose to introduce and overview different aspects of artificial social intelligence and their relationship with LLMs by introducing scientific methods for evaluating social intelligence in LLM.
Outcome: This tutorial will introduce scientific methods for evaluating social intelligence in LLMs, highlighting the key challenges, and identifying promising research directions.
Garbage In, Reasoning Out? Why Benchmark Scores are Unreliable and What to Do About It (2026.findings-eacl)

Copied to clipboard

Challenge: Using social reasoning benchmarks, we uncover pervasive flaws in both benchmark items and evaluation methodology.
Approach: They audit three widely used social reasoning benchmarks and identify flaws in their design and evaluation methodology.
Outcome: The results challenge the validity of current benchmark-based claims about social reasoning in large language models.
A Notion of Complexity for Theory of Mind via Discrete World Models (2024.findings-emnlp)

Copied to clipboard

Challenge: Theory of Mind (ToM) can be used to assess the capabilities of Large Language Models (LLMs) in complex scenarios where social reasoning is required.
Approach: They propose a framework inspired by cognitive load theory to measure the complexity of ToM tasks by a prompting technique that augments the information available to a model with a description of how the environment changes with the agents’ interactions.
Outcome: The proposed framework assesses the complexity of five widely adopted ToM benchmarks and shows that it performs better than other frameworks.
CRoW: Benchmarking Commonsense Reasoning in Real-World Tasks (2023.emnlp-main)

Copied to clipboard

Challenge: Recent efforts in natural language processing (NLP) commonsense reasoning research have produced a number of new datasets and benchmarks.
Approach: They propose a manually-curated, multi-task benchmark that evaluates models' ability to apply commonsense reasoning in the context of six real-world NLP tasks.
Outcome: The proposed benchmark evaluates the ability of models to apply commonsense reasoning in the context of six real-world NLP tasks.
Accommodation and Epistemic Vigilance: A Pragmatic Account of Why LLMs Fail to Challenge Harmful Beliefs (2026.acl-long)

Copied to clipboard

Challenge: Recent studies show that large language models fail to challenge users’ harmful beliefs in domains ranging from medical advice to social reasoning.
Approach: They propose to examine whether pragmatic factors influence LLM accommodation and epistemic vigilance in humans.
Outcome: The proposed model can be understood and addressed as having excessive accommodation and insufficient epistemic vigilance.
Are they lovers or friends? Evaluating LLMs’ Social Reasoning in English and Korean Dialogues (2026.acl-long)

Copied to clipboard

Challenge: Existing studies on LLMs' ability to infer social relationships have limited results for Korean and English.
Approach: They propose a social reasoning task based on a 1.1k-dialogue dataset in English and Korean sourced from movie scripts to evaluate LLMs' ability to infer the social relationships between speakers.
Outcome: The proposed task evaluates the ability of LLMs to infer the social relationships between speakers in 1.1k-dialogue datasets in English and Korean.
VIBE: Can a VLM Read the Room? (2025.findings-emnlp)

Copied to clipboard

Challenge: Vision Language Models (LLMs) cannot account for the role that non-verbal cues play in understanding social situations.
Approach: They propose a task to test the capabilities of Vision Language Models (VLMs) to account for the visual social-pragmatic inference gap.
Outcome: The proposed task tests the capabilities of a VLM for a social reasoning task.
Social Genome: Grounded Social Reasoning Abilities of Multimodal Models (2025.emnlp-main)

Copied to clipboard

Challenge: Social reasoning is a core competency of social intelligence and requires specialized neural and cognitive systems to be able to interpret multimodal interactions.
Approach: They propose to use social reasoning traces to generate fine-grained explanations using external knowledge.
Outcome: The proposed model is based on 272 videos of human interactions and 1,486 human-annotated reasoning traces related to inferences about these interactions.
Bayesian Social Deduction with Graph-Informed Language Models (2026.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) have demonstrated remarkable general-purpose reasoning capabilities across a wide range of tasks.
Approach: They propose a hybrid reasoning framework that externalizes belief inference to a structured probabilistic model while using an LLM for language understanding and interaction.
Outcome: The proposed framework achieves competitive performance with larger models in Agent-Agent play and is the first language agent to defeat human players in a controlled study.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations